Papers by Carlos Daniel Hernandez Mena
SamróMur MilljóN: An ASR Corpus of One Million Verified Read Prompts in Icelandic (2024.lrec-main)
Copied to clipboard
| Challenge: | samrómur is a crowdsourcing web application designed to collect speech data for the advancement of language technologies in Icelandic. |
| Approach: | They propose to use a crowdsourcing web application to collect and verify Icelandic speech data for automatic speech recognition (ASR) they introduce a dataset comprising one million audio clips from the application . |
| Outcome: | The proposed system can produce high-quality speech data for Icelandic . the proposed system is based on a crowdsourced web application built on Mozilla's Common Voice . |
MASRI-HEADSET: A Maltese Corpus for Speech Recognition (2020.lrec-1)
Copied to clipboard
Carlos Daniel Hernandez Mena, Albert Gatt, Andrea DeMarco, Claudia Borg, Lonneke van der Plas, Amanda Muscat, Ian Padovani
| Challenge: | Maltese is the national language of Malta and is spoken by approximately 500,000 people. |
| Approach: | They present the first spoken Maltese corpus designed purposely for Automatic Speech Recognition (ASR) it consists of 8 hours of speech paired with text, recorded by using short text snippets in a laboratory environment. |
| Outcome: | The MASRI-HEADSET corpus was developed by the MASR project at the University of Malta. |
Samrómur Children: An Icelandic Speech Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Samrómur Children contains 131 hours of read speech from Icelandic children aged between 4 to 17 years. |
| Approach: | They propose to build a large-scale speech corpus for automatic speech recognition for Icelandic. |
| Outcome: | The corpus contains 131 hours of read speech from Icelandic children aged 4 to 17 years . the goal of the project is to make Icelandic available in language-technology applications . |